Annals of Internal Medicine
● American College of Physicians
Preprints posted in the last 7 days, ranked by how well they match Annals of Internal Medicine's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Joseph, A.; Kearney, K.; Henricks, C.; Morgan, J. L.; Tan, W.; Shafer, K.; Wrobel, C.; Lacelle, C.; Burns, K.; Jawaid, A.; Tapaskar, N.; Solmonson, A.; Nelson, D. B.; Truby, L. K.
Show abstract
Background: Adult congenital heart disease (ACHD) patients are prone to HLA-antibody formation from multiple surgeries, transfusions, and prosthetic surgical material. Females with ACHD may accrue additional, non-surgical alloantigen exposure. Whether sex modifies the impact of allosensitization on heart transplant (HT) access and outcomes in ACHD remains unknown. Methods: We retrospectively analyzed the OPTN/UNOS registry of adults with ACHD listed for first-time HT (2018-2025). Sensitization was defined by calculated panel reactive antibodies (cPRA) at listing. We tested the sex x sensitization (highly sensitized, cPRA >50%) interaction on transplant access using Fine-Gray competing-risks regression, treating transplantation as the event of interest and death or removal from the waitlist as competing events, and on post-transplant survival using multivariable Cox proportional-hazards regression, both adjusted for age at listing, mechanical support at listing, and the number of distinct prior cardiac surgery categories. Results: Among 856 candidates (38% female), females were more often highly sensitized than males (23% vs 14%; age-adjusted OR 1.81, 95% CI 1.26-2.61), even after adjusting for surgical burden. Sensitization reduced transplant access in females (84% to 71%; median wait 60 to 110 days, p < 0.001) but not males (79% vs 79%, median wait 88 vs 98 days). In adjusted Fine-Gray models, the subdistribution hazard for transplant was reduced in sensitized females (sHR 0.54, 95% CI 0.41-0.72) with no effect in males (sHR 0.96, 95% CI 0.73-1.26), and the sex x sensitization interaction was significant (interaction sHR 0.64, 95% CI 0.44-0.94, p = 0.02). Post-transplant mortality was numerically higher in sensitized than non-sensitized candidates in both sexes and the sex x sensitization interaction on 1-year mortality was not significant. The sex-asymmetric effect persisted and was more pronounced in the multiorgan candidates. Conclusions: Allosensitization is not a sex-neutral barrier to transplant in HT candidates with ACHD. Females are more sensitized and have reduced transplant access without differences in 1-year mortality. The female excess in sensitization is not accounted for by surgical burden, and the exposures responsible remain to be defined. These findings warrant a sex-aware listing strategy and further studies.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Schultz, A. A.; Lange, M.; Shelton, B.; Meinholz, E.; Esselman, D.; Paulsen, E.; Haban, A.; Kesner, V.; Rowe, M.; Burke, R.; Tisler, C.; Tomasallo, C.
Show abstract
Background: Population-based biomonitoring of contemporary-use pesticides remains limited in the United States, particularly in rural agricultural regions, and few studies have repeated measurements within the same individuals over time. Methods: We analyzed 28 urinary pesticide-related biomarkers among 600 adults from the population-based Survey of the Health of Wisconsin with archived urine collected during 2008-2016; 296 participants provided repeat urine and updated exposure information in 2025. Detection frequencies, co-detection, and within-person detection patterns were characterized. Generalized estimating equations were used for stacked, repeated-measures analyses of factors associated with detection of aminomethylphosphonic acid (AMPA), glyphosate, 2,4-dichlorophenoxyacetic acid (2,4-D), and any of these three. Prospective-only analyses evaluated more detailed agricultural and recent exposure measures. Results: Glyphosate, AMPA, and 2,4-D were detected in 7.7%, 6.2%, and 4.3% of retrospective specimens and 5.4%, 3.1%, and 4.1% of prospective specimens, respectively. Co-detection and persistent detection across the 9 to 17-year interval was rare. In repeated-measures models, greater fruit and vegetable intake, older age, and male sex were associated with higher 2,4-D detection. Lower household income was associated with lower AMPA detection, while afternoon/evening collection was associated with higher AMPA detection. In prospective analyses, working on field-crop agricultural land showed the strongest agricultural associations, particularly for 2,4-D and detection of any of the three pesticides. Associations were not seen with self-reported conventional versus organic produce consumption. Conclusions: Urinary pesticide detections were generally infrequent in this Wisconsin population. Diet and direct agricultural activities may be more informative exposure pathways than residing near cropland or private well drinking-water characteristics.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Mathew, Z.; Mehta, R.; Kim, S.; Jeyaraj, J.; Asif, T.
Show abstract
Background: Primary malignant cardiac tumors (PMCTs) are rare and histologically heterogeneous. Objective: To compare demographics, specific ICD-O-3 morphologies, first-course treatment patterns, annual registered case counts, and unadjusted overall survival between soft-tissue and hematologic PMCTs. Methods: We identified 730 PMCT cases diagnosed from 2000 to 2021 in SEER 18 (ICD-O-3 topography C38.0). Histologic lineage was assigned from ICD-O-3 morphology. Comparative analyses included soft-tissue (n=458) and hematologic (n=212) tumors. First-course variables were primary-site surgery, chemotherapy (yes versus no/unknown), and radiotherapy (radiation versus none/unknown). Groups were compared with chi-square tests. Overall survival was estimated with Kaplan-Meier methods; follow-up was truncated at 120 months. Results: Soft-tissue PMCTs occurred predominantly at ages 45-64 years (67.9%), whereas hematologic PMCTs occurred predominantly at age [≥]65 years (63.2%; p<0.001). Men comprised 59.9% of hematologic and 49.3% of soft-tissue cases (p=0.014). The leading soft-tissue morphology was hemangiosarcoma/angiosarcoma (ICD-O-3 9120/3; 201/458, 43.9%); synovial sarcoma accounted for 20/458 cases (4.4%). Diffuse large B-cell lymphoma, NOS, accounted for 131/212 hematologic tumors (61.8%). Any primary-site surgery was recorded in 66.6% of soft-tissue versus 15.6% of hematologic cases (p<0.001). Chemotherapy was recorded in 67.5% versus 51.1% (p<0.001), and radiotherapy in 9.0% versus 20.5% (p<0.001). In exploratory Kaplan-Meier analyses, hematologic patients with recorded chemotherapy had higher unadjusted 120-month overall survival than those without recorded chemotherapy (42.0% versus 12.2%; log-rank p=7.5x10-). Radiation-associated survival differences were not statistically significant in either lineage. Conclusions: Soft-tissue and hematologic PMCTs have distinct age distributions, named histologies, and first-course treatment patterns in SEER. These findings describe registry coding and do not establish treatment effectiveness or population incidence.
Shachar, E. K.; Haas, R.; Rodriguez, V. E.; Lester, J.; Siavoshi, M. A.; Kwan, L.; Niell-Swiller, M.; Spellman, P. T.; Boutros, P. C.; Chang, V. Y.; Karlan, B. Y.
Show abstract
Importance: Chronic stress may contribute to adverse health outcomes through cumulative physiologic dysregulation. Allostatic load (AL), a composite measure of multisystem physiologic burden, may capture biologic effects of structural, social, and psychosocial stress not reflected by self-reported measures. Objective: To evaluate racial and ethnic differences in AL among women with familial cancer risk and examine how socioeconomic status, psychosocial factors, clinical characteristics, and health behaviors contribute to variations in AL. Design: Cross-sectional study of underrepresented minority participants enrolled in the HERSTORY cohort from October 2023 through September 2025, with comparison participants from the UCLA ATLAS biobank. Setting: UCLA academic health system. Participants: The study included 303 racially and ethnically diverse female HERSTORY participants aged [≥]35 years with a family history of cancer and matched non-Hispanic White female ATLAS participants (n=709). Exposures: Race and ethnicity, age, neighborhood deprivation, cancer history and stage, depression, perceived stress, cancer worry, and physical activity. Main Outcomes and Measures: The primary outcome was AL, calculated from cardiometabolic and organ-function measures. A secondary index incorporated race- and ethnicity-specific neutrophil-to-lymphocyte ratio (NLR) derived from 326,826 women in the UCLA Health population. Multivariable regression models evaluated factors associated with elevated AL. Results: Compared with matched non-Hispanic White participants, Black and Asian/Pacific Islander HERSTORY participants had significantly higher AL after adjustment. Hispanic/Latina participants did not have significantly elevated AL. Older age, greater area-level socioeconomic deprivation, and depression were independently associated with higher AL. Prior cancer diagnosis, cancer worry and perceived stress were not significantly associated with AL, whereas regular physical activity was associated with lower AL. Among cancer patients, advanced stage was associated with greater AL. Conclusions and Relevance: This study demonstrates elevated AL among understudied racial/ethnic minority groups with familial cancer risk and identifies associations with neighborhood deprivation, depression, and physical activity. The association between cancer stage and AL suggests that physiologic stress may reflect variation in cancer burden. The lack of association with perceived stress and cancer worry further indicates that physiologic and self-reported psychosocial measures capture distinct dimensions of stress. The development of race/ethnicity-specific NLR thresholds derived from large population samples provide a benchmark for future studies.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Shi, J.; Gu, Q.; Pan, J.; Yang, A.; Fan, M.
Show abstract
Human deep-space missions face bone-kidney risks that cannot be extrapolated from six-month ISS data. We built a 12-state Ca-bone-urine-stone mechanistic ODE model and jointly calibrated its 11 physiological parameters on eight ISS targets by Bayesian identification (M0 base = 19-D; M1 extension adds a GCR-bone coupling term for parsimony testing only), then propagated the M0 posterior to four environments (ISS, Lunar subsurface, Lunar surface, Mars). Lumbar-lower BMD loss increases with mission duration and partial-gravity unloading (ISS 180 d -4.83% -> Mars 730 d -12.15%; 2^3 factorial: duration 82.9%, gravity 12.5%, GCR main effect ~ 0), whereas stone rate follows the opposite gradient (ISS 16.1 vs Mars 13.1 per 1000 person-years), reflecting weakened partial-gravity bone resorption alongside residual urinary chemistry changes. The dominant pathway thus shifts from bone-centric on the ISS to kidney-centric on Mars, where residual urinary-chemistry changes-not bone resorption-drive stone risk. The direct GCR-bone coupling term is unidentifiable at current ISS doses (DeltaWAIC = +0.0076 +/- 0.126 SE), so M0 is retained as the main inference model. Bisphosphonates provide >=84% BMD protection but leave a urinary-chemistry residual, so bisphosphonate monotherapy would underestimate Mars stone risk; potassium-magnesium-citrate combinations (RRR_RSS 51%) should therefore be added to deep-space countermeasures. A Lunar-surface 365-day mission is the earliest environment on the NASA roadmap to cross a composite RED threshold. That profile differs from the regolith-shielded 180-day case in both cumulative GCR (~69x) and duration (2x), so a shielding-specific effect cannot be isolated here; forcing the GCR coupling terms to zero leaves all four composite tiers unchanged (0/4, Supp S24), and the shielded 180-day profile is YELLOW rather than GREEN. Independent hold-out validation (Culliton 2025 60-day HDT-bedrest RCT, n=8 control arm of n=24 total) supports the M0 posterior predictive distribution on the lumbar-BMD sub-scope.
Bingham, J. C.; Arussy, N.
Show abstract
Active Feature Acquisition (AFA) adaptively selects which diagnostic test to order next and offers a route to reduce unnecessary laboratory testing in acute care. Existing clinical AFA evaluations, however, assume every feature can be retrieved on demand and split data at the visit level, both of which inflate apparent performance. We re-evaluate cost-aware AFA under constraints designed to reflect deployment. From MIMIC-IV we constructed a cohort of 64,766 acute admissions (39,884 patients; 21 conditions; 55 features in 30 test panels) with a patient-level split, a 12-hour decision cutoff, and a per-patient availability mask from what was actually measured, and priced panels using the 2026 Medicare fee schedule under panel-level billing. We evaluated EIG-Cost, which scores each panel by Monte-Carlo Expected Information Gain penalised by its dollar cost, against eight published methods across budgets \30--$60 over five patient-level resamples. At a $30 budget, EIG-Cost achieved the highest macro-F1 (0.188, 95% CI [0.185, 0.191]) at the lowest cost ($17.28), exceeding the strongest baseline in all five resamples (p<0.001; Cohen's d=4.0), and led at every budget. Three of the eight methods collapsed to a vitals-only baseline (macro-F1 approx 0.040), acquiring nothing even at higher budgets, a genuine failure to adapt to availability rather than a budget limitation. Despite modest absolute accuracy, EIG-Cost's probabilities were well-calibrated (expected calibration error $0.048$). Under realistic availability constraints, clinical AFA is substantially harder than full-availability benchmarks imply, several published methods fail outright, and cost-aware information-gain scoring is a robust choice in this harder setting.
Ghuman, D.; Achar, T.; Gambhirrao, D.
Show abstract
Background Alcohol-associated injury is a leading cause of emergency department (ED) utilization in the United States and a clinically important driver of preventable morbidity across the adult lifespan. Prior surveillance research has characterized how the rate and severity of alcohol-associated injury vary by patient age, but whether the seasonal timing of injury risk is equally predictable across age groups (a question directly relevant to the timing of clinical screening intensification and public health intervention) has not been formally tested. Methods We conducted a retrospective surveillance analysis of 45,876 alcohol-associated ED visits among adults aged 18 years and older, identified from the National Electronic Injury Surveillance System (NEISS), 2019-2025 (weighted national estimate: 2,092,319 visits), using the structured Alcohol_Involved indicator introduced into NEISS case abstraction in 2019. Patients were stratified by sex and five age groups (18-24, 25-34, 35-49, 50-64, and [≥]65 years). Single-harmonic cosinor (Poisson) regression was used to estimate the seasonal peak day of injury risk (acrophase) for each stratum. To assess reliability, we performed leave-one-year-out jackknife resampling (seven iterations per group), case-resampling bootstrap confidence intervals (1,000 iterations), and likelihood-ratio tests of seasonal-phase interactions. Results Peak injury timing differed significantly across age groups (X^2 [8] = 2356.2, p < .0001). Adults aged 25-64 years showed a highly reproducible early-to-mid-July peak, with jackknife estimates shifting [≤]14 days when any single study year was excluded. Adults aged [≥]65 years showed significant seasonal variation annually (all p < .0001, amplitude comparable to younger groups) but a pooled peak estimate that shifted by up to 100 days across jackknife iterations. Sex-stratified analyses revealed that this instability was driven entirely by females aged [≥]65 years (jackknife range: 332 days, peak consistently in late October through early January) rather than males aged [≥]65 (jackknife range: 31 days, peak consistently in early August). Hospital admission rates increased monotonically with age from 9.0% (18-24 years) to 31.8% ([≥]65 years). Conclusions Alcohol-associated injury follows a reproducible, calendar-stable summer seasonal pattern in adults aged 25-64 years. Among adults [≥]65 years, the previously reported temporal instability is concentrated in the female subgroup, whose seasonal injury risk does not converge on a fixed calendar window. These findings suggest that fixed-calendar prevention and screening strategies are well suited to working-age adults and older men, but older women may require a year-round, individually tailored approach. Keywords: Alcohol-related injury; Emergency department; Seasonality; Age factors; Sex differences; Injury surveillance; Cosinor analysis; Older adults
Gill, P. A.; Bradbury, L. R.; Wang, A.; Hogg, J.; Demase, K.; McKenzie, J.; Fryer, H. A.; Geers, D.; Zaeck, L. M.; Boo, I.; Hogarth, M. P.; Drummer, H. E.; de Vries, R. D.; O'Hehir, R. E.; Sparrow, M. P.; van Zelm, M. C.
Show abstract
Background: Patients receiving anti-TNF treatment for chronic inflammatory disease display impaired antibody responses, but it remains unclear how immune memory formation is affected. We evaluated antibody responses and memory B cells (Bmem) after COVID-19 booster vaccination in inflammatory bowel disease (IBD) patients receiving anti-TNF treatment. Methodology: Blood was sampled at baseline, 1, and 6 months after WH1/BA.5 bivalent or XBB.1.5 monovalent vaccination from 27 IBD patients receiving intravenous anti-TNF and 44 controls. Neutralizing antibodies were measured using an infectious virus assay. SARS-CoV-2 spike receptor binding domain (RBD)-specific serum IgG was quantified by ELISA, and RBD-specific Bmem were immunophenotyped by flow cytometry using recombinant proteins from ancestral, Omicron BA.1, BA.5, XBB.1.5, and JN.1 variants. Results: Serum IgG to vaccine RBD and neutralizing antibodies in patients increased pre to 1 month post-vaccination, but were lower than controls. Ancestral-, BA.5- and XBB.1.5-specific Bmem increased after vaccination but were significantly lower in patients than controls. Within RBD-specific Bmem, frequencies of recently activated CD21lo cells were increased after vaccination, and were higher in patients than controls. Fewer antigen-specific Bmem in patients expressed IgG4, and more expressed IgG3 or IgD following vaccination. Following vaccination, more RBD-specific Bmem recognized multiple viral variants. However, patients had fewer Bmem that could bind to subvariants than controls. Conclusion: Antibody and Bmem responses to COVID-19 booster vaccination in anti-TNF-treated IBD patients displayed reduced capacity, durability and cross-reactivity, suggesting impaired immune memory for protection against breakthrough infection. This supports the recommendation for annual booster vaccination to prevent severe disease and viral spread.
Kim, S. S.; Zissette, S. Z.; Van Meter, C.; Shiiba, M.; Bruck, M.; Tippett, A.; Kamidani, S.; Benkeser, D.; McQuade, E. R.
Show abstract
Importance: Maternal vaccination and long-acting monoclonal antibodies are now available in the U.S. to prevent RSV. Long-acting monoclonal antibody administration in the U.S. commonly occurs after hospital discharge in outpatient settings, leaving some infants unprotected early in life when severe RSV risk is highest. Comparative effectiveness between the two interventions and whether delays affect effectiveness estimates have not been quantified. Objective: To evaluate the effectiveness of infant long-acting monoclonal antibody strategies and a maternal vaccination strategy, each compared to no intervention, and the comparative effectiveness of intervention strategies when accounting for real-world delays in monoclonal antibody receipt. Design: Cohort study using target trial emulation to compare four strategies for prevention of RSV-related outcomes. Setting: The U.S. between 2023 and 2025 using a nationwide database of employer-sponsored commercial insurance claims. Participants: 120,586 commercially insured mother-infants, whose infants were born in the U.S. during the 2023-2024 or 2024-2025 RSV season. Infants who could not be paired with their mother's record, did not enroll in commercial insurance within 75 days from birth, received palivizumab, and had an implausible birth date were excluded. Interventions: Comparison of four RSV prevention strategies: (i) maternal RSVpreF; (ii) long-acting monoclonal antibody given within the first week of life (mAb as intended); (iii) long-acting monoclonal antibody given within a six-month grace period from birth (mAb within grace period); and (iv) a control. Main outcomes and measures: Effectiveness against first RSV-associated hospitalization and medically-attended RSV illness was summarized using adjusted hazard ratios (aHR) and estimated using an inverse propensity weighting approach, with weights accounting for maternal age, maternal comorbidities affecting pregnancy, obstetric and newborn complications, season, region, and birth timing relative to October 1. A weighted Kaplan Meier estimator was used to estimate strategy-specific cumulative incidence of RSV outcomes over time. Results: In the first five weeks of life, the mAb within grace period strategy doubled the hazard of RSV hospitalization (aHR: 2.0 [95% CI: 1.0-4.9]) and increased the hazard of medically-attended RSV (aHR: 1.6 [95% CI: 1.0-2.7]) compared to the maternal RSVpreF strategy. The hazard for RSV hospitalization was similar for the mAb as intended strategy compared to the maternal RSVpreF strategy (aHR = 0.9 [95% CI: 0.3-1.9]). Conclusions and relevance: RSVpreF and monoclonal antibodies were similarly effective when monoclonal antibodies were administered close to birth, but when accounting for real-world delays in monoclonal antibody receipt, the maternal RSVpreF strategy was more effective than the mAb within grace period strategy.
Honore, A.; Rech, T.; Scrivens, A.; Binotto, I.; Zandvoort, C. S.; van der Staaij, H.; Peck, M.; Zivanovic, S.; Stanworth, S. J.; Hartley, C.; Dame, C.; Deschmann, E.; the Neonatal Transfusion Network,
Show abstract
Background and Objectives: Preterm infants are commonly transfused, yet direct cardiorespiratory effects of red blood cell (RBC) transfusions remain poorly understood. We explored the feasibility of using multicentre electronic health data (EHD) to study such cardiorespiratory responses. Methods: Highly granular routine EHD were collected from preterm infants born <32 weeks gestational age at three European centres. Heart rate, oxygen saturation, and respiratory rate were evaluated 12 hours before and after the RBC transfusion. Results: A total of 321 transfusions in 164 infants were analysed. Overall, there was no significant change in the rate of bradycardia and apnoea following transfusion. Cardiorespiratory parameters varied substantially between infants; e.g. 20% of transfusions were associated with an unexpected, significant increase in heart rate. Respiratory rate and oxygen saturation exhibited similarly heterogenous patterns following transfusion. In sub-group analysis, the proportion of transfusions with increased heart rate was significantly higher within the first two weeks than later (32% vs 13%, p=0.0019). Conclusions: Multicentre EHD extraction allows to identify otherwise masked short-term effects of RBC transfusions on cardiorespiratory parameters, possibly indicating cardiac or pulmonary overload. Such effects may vary with adaptation to anaemia. Analysing EHD may ultimately enable personalized transfusion practice.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Watts, K.; Lin, R. C.; Lynch, S.; Warning, J.; Barr, J. J.; Ben Zakour, N.; Campbell, A.; Chan, J.; Collie, L.; Hedges, M.; Hudson, B.; Irwin, A.; Khatami, A.; Kicic, A.; Laucirica, D.; Lauter, C.; Ling, K.-m.; Ng, R.; Pavuk, N.; Rahmatullah, R.; Sinclair, H.; Tucker, E.; Vreugde, S.; Warner, M.; Velickovic, Z.; iredell, j.
Show abstract
Objective As antimicrobial resistance (AMR) continues to threaten global public health, bacteriophage therapy products (BTPs) offer a promising alternative to conventional antimicrobials. However, translation into routine clinical practice requires best practice standards for manufacturing and quality control to ensure the consistent safety, quality, and reliability of personalised BTPs produced for individual patients or small cohorts. Design A modified Delphi methodology was used to develop consensus statements, engaging experts from Australia's National Bacteriophage Therapy Regulatory Working Group across the fields of clinical microbiology, phage biology, good manufacturing practice (GMP), regulatory science, and government. The process comprised three iterative phases: (1) structured statement development, (2) an anonymous REDCap survey, and (3) a hybrid consensus meeting. The strength of evidence and recommendations was assessed using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework. Results Consensus was reached on 35 statements to provide best practice manufacture and quality control guidance for BTPs. These statements address requirements for phage identification and characterisation; define the point at which GMP-aligned processes commence for ubiquitous phages; outline quality control expectations for phage active pharmaceutical ingredient (pAPI) production and maintenance of BTP and host cell repositories. Additional guidance covers quality management systems, including documentation, traceability, and governance. Conclusion These consensus statements provide comprehensive best practice recommendations for the manufacture and quality control of BTPs in Australia. By promoting consistent, safe, and quality-assured approaches to personalised BTPs, they aim to facilitate clinical implementation while remaining aligned with existing international pharmacopoeial standards and regulatory frameworks.